Papers with proxy metric
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets (2026.eacl-long)
Copied to clipboard
| Challenge: | Current large language models (LLMs) are powerful generators of synthetic data, which are used for training smaller, more efficient models. |
| Approach: | They propose a proxy metric for selecting the best LLM generator without human annotations and a metric that measures the performance of a model. |
| Outcome: | The proposed proxy metric outperforms intrinsic heuristics and comes within 0.76 percentage points of the optimal generator baseline. |
The Woman Worked as a Babysitter: On Biases in Language Generation (D19-1)
Copied to clipboard
| Challenge: | a systematic study of biases in natural language generation (NLG) is presented . a study of language models in NLG is conducted by examining language models. |
| Approach: | They propose a systematic study of biases in natural language generation by analyzing text generated from prompts that contain mentions of different demographic groups. |
| Outcome: | The proposed method reveals biases in natural language generation (NLG) by analyzing text generated from demographic prompts. |